Papers with distributional model

5 papers
Can a Gorilla Ride a Camel? Learning Semantic Plausibility from Text (D19-60)

Copied to clipboard

Challenge: Existing work on modeling semantic plausibility has focused on physical plausability but distributional methods fail when tested in supervised settings.
Approach: They propose to use large pretrained language models to model plausibility in supervised settings by extracting attested events from a large corpus and injecting explicit commonsense knowledge into a distributional model.
Outcome: The proposed model is effective in modeling plausibility in a supervised setting.
Building a Web-Scale Dependency-Parsed Corpus from CommonCrawl (L18-1)

Copied to clipboard

Challenge: DepCC is the largest-to-date linguistically analyzed corpus in English . large corpora are essential for the modern data-driven approaches to natural language processing .
Approach: They present a large-to-date linguistically analyzed corpus in English with 365 million documents . they build an index of all sentences and their linguistic meta-data enabling quick search across the corpus .
Outcome: The proposed model outperforms state-of-the-art models on smaller corpora on the SimVerb3500 dataset.
Finely Tuned, 2 Billion Token Based Word Embeddings for Portuguese (L18-1)

Copied to clipboard

Challenge: A distributional semantics model is instrumental to improve the performance of many applications and processing tasks for any language.
Approach: They propose to develop an advanced distributional model for Portuguese with the largest vocabulary and best evaluation scores published so far.
Outcome: The proposed model has the largest vocabulary and the best evaluation scores published so far.
When Hearst Is not Enough: Improving Hypernymy Detection from Corpus with Distributional Models (2020.emnlp-main)

Copied to clipboard

Challenge: a taxonomy is a semantic hierarchy of words or concepts organized w.r.t. their hypernymy relationships.
Approach: They propose a framework for hypernymy detection using large textual corpora . they quantify the non-negligible existence of specific sparsity cases .
Outcome: The proposed framework quantifies the non-negligible existence of specific sparsity cases on several benchmark datasets.
A Multi-word Expression Dataset for Swedish (2020.lrec-1)

Copied to clipboard

Challenge: Existing data on compositionality of multi-word expressions is limited and only available for high resource languages.
Approach: They present a set of Swedish multi-word expressions annotated with degree of compositionality . they also consider syntactically complex constructions and publish a formal specification of each expression .
Outcome: The proposed dataset includes 96 Swedish multi-word expressions with degree of compositionality.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations